Papers with speech enhancement
Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model (2024.emnlp-main)
Copied to clipboard
Xiangyu Zhang, Daijiao Liu, Hexin Liu, Qiquan Zhang, Hanyu Meng, Leibny Paola Garcia Perera, EngSiong Chng, Lina Yao
| Challenge: | Existing approaches to enhance inference speed and training require complex modifications to the model. |
| Approach: | They propose to double the training and inference speed of Denoising Diffusion Probabilistic Models by simply redirecting the generative target to the wavelet domain. |
| Outcome: | The proposed method doubles the training and inference speed of Speech DDPMs by redirecting the generative target to the wavelet domain. |
Far-Field Speaker Recognition Benchmark Derived From The DiPCo Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Using a publicly-available corpus, we propose a far-field speaker verification benchmark. |
| Approach: | They propose a far-field speaker verification benchmark derived from the publicly available DiPCo corpus. |
| Outcome: | The proposed tasks are very challenging and hope to inspire the speech community to develop new methods and systems for this challenging domain. |
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing (2022.acl-long)
Copied to clipboard
Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei
| Challenge: | Existing work shows that pre-trained models can improve in various natural language processing tasks. |
| Approach: | They propose a unified-modal encoder-decoder framework that pre-trains speech-text representations using large-scale unlabeled speech and text data. |
| Outcome: | The proposed framework is superior to existing models on speech-to-text processing tasks. |
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement (2025.acl-long)
Copied to clipboard
Boyi Kang, Xinfa Zhu, Zihan Zhang, Zhen Ye, Mingshuai Liu, Ziqian Wang, Yike Zhu, Guobin Ma, Jun Chen, Longshuai Xiao, Chao Weng, Wei Xue, Lei Xie
| Challenge: | Recent advances in language models have demonstrated strong capabilities in semantic understanding and contextual modeling. |
| Approach: | They propose a LLaMA-based language model that incentivizes generalization capabilities for speech enhancement. |
| Outcome: | The proposed language model outperforms prior task-specific discriminative and generative models in acoustic enhancement tasks. |